Search CORE

3 research outputs found

Revisiting the challenges and surveys in text similarity matching and detection methods

Author: Kusrini Kusrini
Muhammad Alva Hendi
Oyong Irwan
Publication venue: 'Universitas Ahmad Dahlan, Kampus 3'
Publication date: 30/09/2022
Field of study

The massive amount of information from the internet has revolutionized the field of natural language processing. One of the challenges was estimating the similarity between texts. This has been an open research problem although various studies have proposed new methods over the years. This paper surveyed and traced the primary studies in the field of text similarity. The aim was to give a broad overview of existing issues, applications, and methods of text similarity research. This paper identified four issues and several applications of text similarity matching. It classified current studies based on intrinsic, extrinsic, and hybrid approaches. Then, we identified the methods and classified them into lexical-similarity, syntactic-similarity, semantic-similarity, structural-similarity, and hybrid. Furthermore, this study also analyzed and discussed method improvement, current limitations, and open challenges on this topic for future research directions

Journal of Education and Learning (EduLearn)

Revisiting the challenges and surveys in text similarity matching and detection methods

Author: Kusrini Kusrini
Muhammad Alva Hendi
Oyong Irwan
Publication venue: 'Universitas Ahmad Dahlan, Kampus 3'
Publication date: 30/09/2022
Field of study

Journal of Education and Learning (EduLearn)

UAD Journal Management System

Arabic Language Opinion Mining Based on Long Short-Term Memory (LSTM)

Author: Abdullah Alomair
Arief Setyanto
Arif Laksito
Fawaz Alarfaj
Irwan Oyong
Kusrini
Lilis Kurniasari
Mardhiya Hayaty
Mohammed Alreshoodi
Naif Almusallam
Publication venue: 'MDPI AG'
Publication date: 20/04/2022
Field of study

Arabic is one of the official languages recognized by the United Nations (UN) and is widely used in the middle east, and parts of Asia, Africa, and other countries. Social media activity currently dominates the textual communication on the Internet and potentially represents people’s views about specific issues. Opinion mining is an important task for understanding public opinion polarity towards an issue. Understanding public opinion leads to better decisions in many fields, such as public services and business. Language background plays a vital role in understanding opinion polarity. Variation is not only due to the vocabulary but also cultural background. The sentence is a time series signal; therefore, sequence gives a significant correlation to the meaning of the text. A recurrent neural network (RNN) is a variant of deep learning where the sequence is considered. Long short-term memory (LSTM) is an implementation of RNN with a particular gate to keep or ignore specific word signals during a sequence of inputs. Text is unstructured data, and it cannot be processed further by a machine unless an algorithm transforms the representation into a readable machine learning format as a vector of numerical values. Transformation algorithms range from the Term Frequency–Inverse Document Frequency (TF-IDF) transform to advanced word embedding. Word embedding methods include GloVe, word2vec, BERT, and fastText. This research experimented with those algorithms to perform vector transformation of the Arabic text dataset. This study implements and compares the GloVe and fastText word embedding algorithms and long short-term memory (LSTM) implemented in single-, double-, and triple-layer architectures. Finally, this research compares their accuracy for opinion mining on an Arabic dataset. It evaluates the proposed algorithm with the ASAD dataset of 55,000 annotated tweets in three classes. The dataset was augmented to achieve equal proportions of positive, negative, and neutral classes. According to the evaluation results, the triple-layer LSTM with fastText word embedding achieved the best testing accuracy, at 90.9%, surpassing all other experimental scenarios

Multidisciplinary Digital Publishing Institute